Repository navigation
feat(infra): training image + GPU Job runbook for ML squad - #30
Merged
Merged
Conversation
Egitim kodu (services/ml/training) bir konteynere girmeden K8s Job olarak kosturulamiyordu; infra/docker/README'de listelenen training.Dockerfile mevcut degildi. - infra/docker/training.Dockerfile: stok pytorch CUDA imaji (torch+cuDNN hazir) uzerine requirements/ml.txt. Yalnizca services/ml kopyalanir, non-root calisir, Hydra ciktilari icin yazilabilir /workspace. - requirements/ml.txt: hydra-core, omegaconf, boto3, python-dotenv eklendi. Egitim giris noktasi bunlari import ediyordu ama listede yoktular; imaj build oluyor, ilk kosu ImportError ile duserdi. - docs/runbooks/training-job.example.yaml: GPU Job sablonu. /dev/shm icin emptyDir — varsayilan 64 MB PyTorch DataLoader worker'larini "Bus error" ile dusuruyor. - docs/runbooks/ML-TRAINING.md: DevOps'un bir kerelik kurulumu (imaj build, MinIO SealedSecret) + ML squad'in kosu dongusu + sorun giderme tablosu. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
feat(infra): training image + GPU Job runbook
The ML squad could not run training on the server:
infra/docker/README.mdlisted
training.Dockerfile, but the file did not exist. Without an image,the training code cannot run as a Kubernetes
Job.Changes
infra/docker/training.Dockerfile— built on the stock PyTorch CUDAimage rather than pip-installing torch (avoids a multi-GB download and the
risk of a CUDA/driver version mismatch). Runs as a non-root user; only
services/mlis copied into the image.infra/docker/training.Dockerfile.dockerignore— limits the buildcontext to
requirements/andservices/ml/, keeping the Go API and theNext.js frontend out of the build.
requirements/ml.txt— addedhydra-core,omegaconf,boto3, andpython-dotenv.services/ml/training/train.pyandservices/ml/minio_loader.pyimport these, but they were missing from thelist: the image would build fine and the first run would fail with
ImportError.docs/runbooks/training-job.example.yaml— GPUJobtemplate. Includesan
emptyDirmounted at/dev/shm; the container default is 64 MB, whichPyTorch
DataLoaderworkers exhaust, crashing withBus error.docs/runbooks/ML-TRAINING.md— one-time DevOps setup (image build andpush, MinIO SealedSecret) plus the ML squad's day-to-day run loop and a
troubleshooting table.
Notes
Python version. The image ships Python 3.11 (stock PyTorch CUDA image)
while the repo root requires 3.13. This divergence is deliberate: th
floor exists for
ehtim(data squad) and the training code does not use it.The rationale is documented in the Dockerfile header.
Image tags. Tags are date-based (
2026-08-07);latestis deliberatelyavoided because Kubernetes will not re-pull an unchanged tag, so an image
update would silently fail to apply.
Question for the ML squad
services/ml/conf/data/default.yamlsetsbucket_name: datasetsandminio_prefix: datasets/training-512/v1, while the layout indocs/ is adatasetsbucket with atraining-512/v1/prefix. If both are correct as written, boto3 will look fordatasets/datasets/training-512/v1nothing. Worth verifying withmc ls dh/datasets/` before the first run.Testing
Not yet built on the server — the training code lives on `ml/feature
is not merged. Build recipe is in the runbook; the Dockerfile's fina
step imports every required package and asserts that torch was built
CUDA, so a missing dependency fails the build rather than the third
training run.
main...devops/featur
PR açıklaması (kopyala)
▎ feat(infra): training image + GPU Job runbook
▎
▎ ML ekibinin sunucuda eğitim koşabilmesi için eksik olan parçalar. ining.Dockerfile'ı listeliyordu ama dosya yoktu — kod bir konteyneregirmeden Job olarak koşturulamıyordu.
▎
▎ - training.Dockerfile — stok PyTorch CUDA imajı tabanlı (torch pip ile kurulmadı: GB'larca indirme + CUDA sürüm eşleşmesi riski). Non-root, sadece services/ml kopyalanıyor.
▎ - requirements/ml.txt — hydra-core, omegaconf, boto3, python-dotenv eklendi. train.py ve minio_loader.py bunları import ediyordu ama listede yoktular; imaj build
▎ olur, ilk koşu ImportError ile düşerdi.
▎ - training-job.example.yaml — GPU Job şablonu. /dev/shm için emptyDir: varsayılan 64 MB, PyTorch DataLoader worker'larını Bus error ile düşürüyor.
▎ - ML-TRAINING.md — DevOps'un bir kerelik kurulumu + ML squad'ın ko.
▎
▎ Not: İmajın Python'u 3.11 (stok PyTorch imajı), repo kökü 3.13 istiyor. Fark bilinçli — 3.13 alt sınırı ehtim için, eğitim kodu onu kullanmıyor. Gerekçe Dockerfile başlığında.
▎
▎ ML squad'a soru: conf/data/default.yaml bucket_name: datasets ve minio_prefix: datasets/training-512/v1 diyor. docs/DATA.md düzeni datasets bucket'ı +
▎ training-512/v1/ prefix'i. İkisi birlikteyse boto3 datasets/datasence doğrulanmalı.
Merge sonrası sunucuda build için origin/devops/feature yerine origin/main'den alabilirsin:
cd ~/deephorizon && git fetch origin
git checkout -B build/training origin/ml/feature
git checkout origin/main -- infra/docker/training.Dockerfile infra/docker/training.Dockerfile.dockerignore requirements/ml.txt
TAG=$(date +%F)
docker build -f infra/docker/training.Dockerfile -t localhost:32000/deephorizon-training:$TAG .
Build çıktısında takılırsan logu at, devops/fix ile düzeltiriz.
✻ Sautéed for 25s
❯ PR açıklamasını ingilizce ve .md formatta yaz
feat(infra): training image + GPU Job runbook
The ML squad could not run training on the server:
infra/docker/REA listedtraining.Dockerfile, but the file did not exist. Without an image, the training code cannot run as a KubernetesJob`.Changes
infra/docker/training.Dockerfile— built on the stock PyTorch CUDAimage rather than pip-installing torch (avoids a multi-GB download and the
risk of a CUDA/driver version mismatch). Runs as a non-root user;
services/mlis copied into the image.infra/docker/training.Dockerfile.dockerignore— limits the bcontext to
requirements/andservices/ml/, keeping the Go API and theNext.js frontend out of the build.
requirements/ml.txt— addedhydra-core,omegaconf,boto3, andpython-dotenv.services/ml/training/train.pyandservices/ml/minio_loader.pyimport these, but they were missing from thelist: the image would build fine and the first run would fail with
ImportError.docs/runbooks/training-job.example.yaml— GPUJobtemplatean
emptyDirmounted at/dev/shm; the container default is 64 MB, whichPyTorch
DataLoaderworkers exhaust, crashing withBus error.docs/runbooks/ML-TRAINING.md— one-time DevOps setup (image build andpush, MinIO SealedSecret) plus the ML squad's day-to-day run loop
troubleshooting table.
Notes
Python version. The image ships Python 3.11 (stock PyTorch CUDA image)
while the repo root requires 3.13. This divergence is deliberate: the 3.13
floor exists for
ehtim(data squad) and the training code does not use it.The rationale is documented in the Dockerfile header.
Image tags. Tags are date-based (
2026-08-07);latestis deliberatelyavoided because Kubernetes will not re-pull an unchanged tag, so an
update would silently fail to apply.
Question for the ML squad
services/ml/conf/data/default.yamlsetsbucket_name: datasetsandminio_prefix: datasets/training-512/v1, while the layout indocs/ is adatasetsbucket with atraining-512/v1/prefix. If both are correct as written, boto3 will look fordatasets/datasets/training-512/v1and find nothing. Worth verifying withmc ls dh/datasets/` before the firstTesting
Not yet built on the server — the training code lives on
ml/feature, whichis not merged. Build recipe is in the runbook; the Dockerfile's final
RUNstep imports every required package and asserts that torch was built
CUDA, so a missing dependency fails the build rather than the third
training run.